G3: Genes, Genomes, Genetics
◐ Oxford University Press (OUP)
All preprints, ranked by how well they match G3: Genes, Genomes, Genetics's content profile, based on 252 papers previously published here. The average preprint has a 0.20% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Trimble, C. J.; Laumer, C. E.; Lamare, M.; Carter, H. F.; Byrne, M.; Fellner, M.; Williams, S. T.; Kenny, N. J.
Show abstract
Linckia laevigata is a tropical sea star commonly found throughout the Indian and Pacific Oceans, and is one of the top ten-most collected invertebrates, often encountered in the aquarium trade. It has been the subject of investigations into population structure, biodiversity and ecology, particularly regarding gene flow among populations throughout its range, the status of different colour morphs and its relationship with its putative sister species, L. multifora. Here we present and describe a high-quality genome assembly for L. laevigata. Our assembly is 585.97 Mb in length, with a scaffold N50 of 3 Mb. The genome has a typical repeat (36.99%) and GC content (41.33%), when compared with other echinoderm datasets. Our genome annotation recovers 16,178 genes, with high (89.4%) recovery of the metazoan BUSCO set. This novel resource will provide a model organism for studying the biogeography of the tropical Indo-West Pacific region, and more specifically facilitate the investigation of a range of sea star traits at the genomic level. SignificanceSea stars remain poorly characterised at a genomic level. Data from the charismatic tropical sea star Linckia laevigata will enable future studies into the biology, population structure and evolution of these ecologically important species.
Ryan, C.; Fraser, F.; Irish, N.; Barker, T.; Knitlhoffer, V.; Durrant, A.; Reynolds, G.; Kaithakottil, G.; Swarbreck, D.; De Vega, J. J.
Show abstract
Haplotyped-resolved phased assemblies aim to capture the full allelic diversity in heterozygous and polyploid species to enable accurate genetic analyses. However, building non-collapsed references still presents a challenge. Here, we used long-range interaction Hi-C reads (high-throughput chromatin conformation capture) and HiFi PacBio reads to assemble the genome of the apomictic cultivar Basilisks from Urochloa decumbens (2n = 4x = 36), an outcrossed tetraploid Paniceae grass widely cropped to feed livestock in the tropics. We identified and removed Hi-C reads between homologous unitigs to facilitate their scaffolding and employed methods for the manual curation of rearrangements and misassemblies. Our final phased assembly included the four haplotypes in 36 chromosomes. We found that 18 chromosomes originated from diploid U. brizantha and the other 18 from either U. ruziziensis or diploid U. decumbens. We also identified a chromosomal translocation between chromosomes 5 and 32, as well as evidence of pairing exclusively within subgenomes, except for a homoeologous exchange in chromosome 21. Our results demonstrate that haplotype-aware assemblies accurately capture the allelic diversity in heterozygous species, making them the preferred option over collapsed-haplotype assemblies.
Bickerstaff, J. R.; Walsh, T.; Court, L.; Pandey, G.; Ireland, K.; Cousins, D.; Caron, V.; Wallenius, T.; Slipinski, A.; Rane, R.; Escalona, H.
Show abstract
Bark and ambrosia beetles are among the most ecologically and economically damaging introduced plant pests worldwide, with life history traits including polyphagy, haplodiploidy, inbreeding polygyny and symbiosis with fungi contributing to their dispersal and impact. Species vary in host tree ecologies, with many attacking stressed or recently dead trees, such as the globally distributed E. similis (Ferrari). Other species, like the Polyphagous Shot Hole Borer (PSHB) Euwallacea fornicatus (Eichhoff), can attack over 680 host plants and is causing considerable economic damage in several countries worldwide. Despite their notoriety, publicly accessible genomic resources for Euwallacea Hopkins species are scarce, hampering better understanding of their invasive capabilities as well as modern control measures, surveillance and management. Using a combination of long and short read sequencing platforms we assembled and annotated high quality (BUSCO > 98% complete) chromosome level genomes for these species. Comparative macro-synteny analysis showed an increased number of chromosomes in the haplodiploid inbreeding species of Euwallacea compared to diploid outbred species, due to fission events. This suggests that life history traits can impact chromosome structure. Further, the genome of E. fornicatus had a higher relative proportion of repetitive elements, up to 17% more, than E. similis. Additionally, metagenomic assembly pipelines identified microbiota associated with both species including Fusarium fungal symbionts and a novel Wolbachia strain. These novel genomes of haplodiploid inbreeding species will contribute to the understanding of how life history traits are related to their evolution and will contribute to the management of these invasive pests. SignificanceScolytinae are significant forestry pests around the world and commonly translocated due to human trade of wood and plant products. Life history traits including inbreeding and haplodiploidy are attributed to their successful establishment in novel environments. Euwallacea fornicatus is widely distributed and attacks a wide variety of live host trees. This study reports the genome of this species and for, E. similis, which colonises dead host trees. The genome assemblies presented herein are highly complete and scaffolded to pseudo-chromosomal level. Comparative analyses of these genomes and of other Scolytinae highlight significant chromosomal rearrangements between haplodiploid inbreeding Euwallacea and diploid outbreeding scolytinae species. Higher relative proportions of transposable elements were identified E. fornicatus, which may promote the species ability to attack live host trees. These genomes are the first for haplodiploid beetles and will contribute to the understanding of evolution of life history traits and the management of invasive insects.
Paul, S.; Stamnes, M. A.; Moye-Rowley, W. S.
Show abstract
Transcriptional regulation of azole resistance in the filamentous fungus Aspergillus fumigatus is a key step in development of this problematic clinical phenotype. We and others have previously described a C2H2-containing transcription factor called FfmA that is required for normal levels of voriconazole susceptibility and expression of an ATP-binding cassette transporter gene called abcG1. Null alleles of ffmA exhibit a strongly compromised growth rate even in the absence of any external stress. Here we employ an acutely repressible doxycycline-off form of ffmA to rapidly deplete FfmA protein from the cell. Using this approach, we carried out RNA-seq analyses to probe the transcriptome of A. fumigatus cells that have been deprived of normal FfmA levels. We found that 2000 genes were differentially expressed upon depletion of FfmA, consistent with the wide-ranging effect of this factor on gene regulation. Chromatin immunoprecipitation coupled with high throughput DNA sequencing analysis (ChIP-seq) identified 530 genes that were bound by FfmA using two different antibodies for immunoprecipitation. More than 300 of these genes were also bound by AtrR demonstrating the striking regulatory overlap with FfmA. However, while AtrR is clearly an upstream activation protein with clear sequence specificity, our data suggest that FfmA is a chromatin-associated factor that may bind to DNA in a manner dependent on other factors. We provide evidence that AtrR and FfmA interact in the cell and can influence one anothers expression. This interaction of AtrR and FfmA is required for normal azole resistance in A. fumigatus.
Pipkin, H. J. J.; Lindsay, H. L.; Smiley, A. T.; Jurmu, J. D.; Arsham, A. M.
Show abstract
The compound eye of Drosophila melanogaster has long been a model for studying genetics, development, neurodegeneration, and heterochromatin. Imaging and morphometry of adult Drosophila and other insects is hampered by the low throughput, narrow focal plane, and small image sensors typical of stereomicroscope cameras. When data collection is distributed among many individuals or extended time periods, these limitations are compounded by inter-operator variability in lighting, sample positioning, focus, and post-acquisition processing. To address these limitations we developed a method for multiplexed quantitative analysis of adult Drosophila melanogaster phenotypes. Efficient data collection and analysis of up to 60 adult flies in a single image with standardized conditions eliminates inter-operator variability and enables precise quantitative comparison of morphology. Semi-automated data analysis using ImageJ and R reduces image manipulations, facilitates reproducibility, and supports emerging automated segmentation methods, as well as a wide range of graphical and statistical tools. These methods also serve as a low-cost hands-on introduction to imaging, data visualization, and statistical analysis for students and trainees.
Eroglu, M.; Hobert, O.
Show abstract
Tagging all proteins encoded by an animal genome with a fluorescent tag would open many windows to the discovery of unexpected patterns of protein expression and localization. To scale such an approach, it would be beneficial to introduce multiple, spectrally distinct fluorophore tags in parallel. As a first step in this direction, we undertook a pilot study in the nematode C. elegans, in which we set out to tag 30 different genetic loci with three different fluorophores, with 3 tags being introduced at a time. By choosing essential genes, predicted based on transcriptomics to cover a range of expression levels, we explore issues relating to disrupting gene function and visibility of tagged proteins. We demonstrate that such a tagging approach is highly efficient and indeed reveals unanticipated patterns of cellular sites of expression, as well as subcellular protein localization. We hope that this pilot study will motivate attempts to scale this tagging approach to more loci and, ultimately, the whole genome.
Thompsky, B.; Beraut, E.; Cooper, R. D.; Escalona, M.; Espinoza, R. E.; Fisher, R. N.; Miller, C.; Nguyen, O.; Sacco, S.; Sahasrabudhe, R.; Seligmann, W. E.; Tofflemier, E.; Wang, I. J.; Shaffer, H. B.
Show abstract
We assembled and annotated a chromosome-level reference genome for the Western Spadefoot, Spea hammondii (Anura, Scaphiopodidae) representing one of only three amphibians included in the California Conservation Genomics Project (CCGP). Spea hammondii is a vernal pool breeding anuran native to California and northwestern Baja California which has undergone both range contractions and local extirpations across its distribution, primarily due to habitat loss and degradation and drought. The species is recognized by the state of California as a Species of Special Concern and is proposed for listing under the United States Endangered Species Act. Using the established CCGP pipeline, this S. hammondii genome was produced using Pacific Biosciences HiFi long-reads and Omni-C proximity ligation, resulting in a de novo genome assembly 1.14 Gb in length, distributed across 479 scaffolds (scaffold N50 = 120.8 Mb; largest scaffold = 183.6 Mb) with a BUSCO completeness score of 90.9% using a conserved tetrapod ortholog set. Our assembly shows high base accuracy (QV = 63.7) and low frameshift error in coding regions (QV 50.42). Annotation of this genome yielded 20,434 genes with a BUSCO completeness score of 94.7%. This reference genome, in combination with range-wide resequencing data from CCGP, will facilitate statewide population genomic assessments to delineate conservation units, quantify inbreeding and genomic load, and test for adaptive variation associated with vernal pool hydrology and drought tolerance, all of which are important considerations in the proposed federal listing.
Lieser, B. C.; Larsen, C. I. S.; Bonthius, R. M.; Thompson, J. S.; Toering Peters, S.
Show abstract
Gene model for the ortholog of Phosphatidylinositol 3-kinase 21B (Pi3K21B) in the D. eugracilis May 2021 (Stanford ASM1815383v1/DeugRefSeq2) Genome Assembly (GenBank Accession: GCF_018153835.1) of Drosophila eugracilis. This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Nolte, N. F.; Petek, M.; Angulo Lara, P.; Nicassio, F.; Marroni, F.; McIntyre, L.
Show abstract
Inaccurate allele and gene expression counts due to map bias and genome ambiguity lead to high false positive and false negative rates in studies of allelic imbalance. We demonstrate that long read RNA-seq and straightforward quality control measures can be used to reduce bias in allele counts in case studies from four species: Drosophila melanogaster, a diploid insect; Solanum tuberosum, an autopolyploid plant; Pongo abelii, a highly heterozygous diploid primate, and Homo sapiens. We recommend 1) mapping to a personalized genome to increase the number of allele assignments 2) tracking multimapping reads and tuning mapping parameters to ensure accurate allele and gene expression counts and 3) evaluating apparent extreme allele bias to identify errors in genome assembly and annotation. We show that these steps can be executed in a straightforward manner and recommend tools for each step.
Mongue, A. J.; Markee, A.; Grebler, E.; Liesenfelt, T.; Powell, E. C.
Show abstract
Scale insects are of interest both to basic researchers for their unique reproductive biology and to applied researchers for their pest status. In spite of this interest, there remain few genomic resources for this group of insects. To begin addressing this lack of data, we present the genome sequence of the tuliptree scale insect, Toumeyella liriodendri (Gmelin) (Hemiptera: Coccomorpha: Coccidae). The genome assembly spans 536Mb, with over 96% of sequence assembled into one of 17 chromosomal scaffolds. We characterize roughly 66% of this sequence as repetitive and annotate 16,508 protein coding genes. Then we use the reference genome to explore the phylogeny of soft scales (Coccidae) and evolution of karyotype within the family. We find that T. liriodendri is an early-diverging soft scale, less closely related to most sequenced soft scales than a species of the family Aclerdidae is. This molecular result bolsters a previous, character-based phylogenetic placement of Aclerdidae within Coccidae. In terms of genome structure, T. liriodendri has nearly twice as many chromosomes as the only other soft scale assembled to the chromosome level, Ericerus pela (Chavannes). In comparing the two, we find that chromosome number evolution can largely be explained by simple fissions rather than more complex rearrangements. These genomic natural history observations lay a foundation for further exploration of this unique group of insects.
Colp, M. J.; Archibald, J. M.
Show abstract
Acanthamoeba castellanii is a free-living amoeba that is emerging as a model organism for the study of eukaryotic microbiology. It is one of the most widely studied members of the Amoebozoa, and is both an important grazer in soil communities and an opportunistic human pathogen; A. castellanii is thus of evolutionary, ecological, and biomedical significance. Despite its potential as a lab workhorse, the genome biology of A. castellanii is complex and poorly understood. Polyploidy is a common feature of many amoebozoan genomes, and members of the genus Acanthamoeba are no exception; they appear to be not only polyploid, where genome copy number is inflated beyond the conventional haploid and diploid states, but also aneuploid, i.e., with inter-chromosomal copy number variation. To better understand aneuploidy in A. castellanii and how it may vary over time and between closely related strains, we analyzed nanopore and Illumina sequence datasets from several wild-type and mutant A. castellanii lines, with a focus on quantifying single nucleotide polymorphism (SNP) and structural variant allele frequencies across chromosome-scale scaffolds. Our findings suggest that intragenomic chromosome copy number is highly variable in Acanthamoeba and can change dynamically even over laboratory time scales. Significance StatementAcanthamoeba castellanii is becoming an important model organism for basic and applied research. However, its apparent polyploidy and aneuploidy has the potential to complicate the interpretation of results that depend on knowledge of gene copy number. In this study, we reveal the complex nature of ploidy in this organism by analyzing long- and short-read sequence data. Our results provide a reference point against which genomic and experimental data from A. castellanii can be interpreted, and guide future efforts aimed at more precisely characterizing how the organism regulates its genome and chromosome copy number.
Lawson, M. E.; Wellik, I. G.; Le, V.; Bennett, E. N.; Thompson, J. S.; Page, S. T.; Rele, C. P.; Hark, A. T.
Show abstract
Gene model for the ortholog of Glycogen phosphorylase (Glyp) in the Apr. 2013 (BCM-HGSC Dpse_3.0/DpseGB3) Genome Assembly of D. pseudoobscura (GCA_000001765.2). This ortholog was characterized as part of a developing dataset to study the evolution of the Insulin/insulin-like growth factor signaling pathway (IIS) across the genus Drosophila using the Genomics Education Partnership gene annotation protocol for Course-based Undergraduate Research Experiences.
Cao, W. X.; Merritt, D.; Pe, K.; Cesar, M.; Hobert, O.
Show abstract
One problem that has hampered the use of red fluorescent proteins in the fast-developing nematode C. elegans has been the substantial time delay in maturation of several generations of red fluorophores. The recently described mScarlet-I3 protein has properties that may overcome this limitation. We compare here the brightness and maturation time of CRISPR/Cas9 genome-engineered mScarlet, mScarlet3, mScarlet-I3 and GFP reporter knock-ins. Comparing the onset and brightness of expression of reporter alleles of C. elegans golg-4, encoding a broadly expressed Golgi resident protein, we found that the onset of detection of mScarlet-I3 in the embryo is several hours earlier than older versions of mScarlet and comparable to GFP. These findings were further supported by comparing mScarlet-I3 and GFP reporter alleles for pks-1, a gene expressed in the CAN neuron and cells of the alimentary system, as well as reporter alleles for the panneuronal, nuclear marker unc-75. Hence, the relative properties of mScarlet-I3 and GFP do not depend on cellular or subcellular context. In all cases, mScarlet-I3 reporters also show improved signal-to-noise ratio compared to GFP.
Cadena, D.; Pabon, L.; DoNascimiento, C.; Abueg, L.; Tiley, T.; O-Toole, B.; Absolon, D.; Sims, Y.; Formenti, G.; Fedrigo, O.; Jarvis, E.; Torres, M.
Show abstract
Animals living in caves are of broad relevance to evolutionary biologists interested in understanding the mechanisms underpinning convergent evolution. In the Eastern Andes of Colombia, populations from at least two distinct clades of Trichomycterus catfishes (Siluriformes) independently colonized cave environments and converged in phenotype by losing their eyes and pigmentation. We are pursuing several research questions using genomics to understand the evolutionary forces and molecular mechanisms responsible for repeated morphological changes in this system. As a foundation for such studies, here we describe a diploid, chromosome-scale, long-read reference genome for Trichomycterus rosablanca, a blind, depigmented species endemic to the karstic system of the department of Santander. The nuclear genome comprises 1Gb in 27 chromosomes, with a 40.0x HiFi long-read genome coverage having a N50 scaffold of 40.4 Mb and N50 contig of 13.1 Mb, with 96.9% (Eukaryota) and 95.4% (Actinopterygii) universal single-copy orthologs (BUSCO). This assembly provides the first reference genome for the speciose genus Trichomycterus, which will serve as a key resource for research on the genomics of phenotypic evolution.
Lew-Smith, J.; Binkley, J.; Sherlock, G.
Show abstract
The Candida Genome Database (CGD; www.candidagenome.org) is unique in being both a model organism database and a fungal pathogen database. As a fungal pathogen database, CGD hosts locus pages for five species of the best-studied pathogenic fungi in the Candida group. As a model organism database, the species Candida albicans serves as a model both for other Candida spp. and for non-Candida fungi that form biofilms and undergo routine morphogenic switching from the planktonic form to the filamentous form, which is not done by other model yeasts. As pathogenic Candida species have become increasingly drug resistant, the high lethality of invasive candidiasis in immunocompromised people is increasingly alarming. There is a pressing need for additional research into basic Candida biology, epidemiology and phylogeny, and potential new antifungals. CGD serves the needs of this diverse research community by curating the entire gene-based Candida experimental literature as it is published, extracting, organizing and standardizing gene annotations. Most recently, we have begun linking clinical data on disease to relevant Literature Topics to improve searchability for clinical researchers. Because CGD curates for multiple species and most research focuses on aspects related to pathogenicity, we focus our curation efforts on assigning Literature Topic tags, collecting detailed mutant phenotype data, and assigning controlled Gene Ontology terms with accompanying evidence codes. Our Summary pages for each feature include the primary name and all aliases for that locus, a description of the gene and/or gene product, detailed ortholog information with links, a JBrowse window with a visual view of the gene on its chromosome, summarized phenotype, Gene Ontology, and sequence information, references cited on the summary page itself, and any locus notes. The database serves as a community hub, where we link to various types of reference material of relevance to Candida researchers, including colleague information, news, and notice of upcoming meetings. We routinely survey the community to learn how the field is evolving and how needs may have changed. A key future challenge is management of the flood of high-throughput expression data to make it as useful as possible to as many researchers as possible. The central challenge for any community database is to turn data into knowledge, which the community can access, use, and build upon.
Maclary, E. T.; Shapiro, M. D.
Show abstract
Plumage pigmentation plays critical roles in survival and reproductive success in birds, from providing camouflage and thermoregulation to mediating elaborate mating displays. The genetic and developmental origins of diverse plumage pigmentation patterns remain incompletely understood in part due to limited intraspecific variation and high levels of genetic divergence between distantly related species. Domestic avian species are more tractable models for understanding the genetic architecture of plumage pigmentation, but the relevance of domestic phenotypes to plumage patterns observed in the wild is not clear. Here, we used comparative genomic approaches to examine coding variation in EDNRB2, a candidate gene associated with loss of plumage melanin in several species, in representative genomes from a diverse array of wild and domestic birds. We found widespread coding variation in EDNRB2 and in other pigmentation genes with limited pleiotropic roles in development. We also found that EDNRB2- mediated melanin loss may play a critical role in establishing bright non-melanin plumage colors. This work highlights EDNRB2 as a key candidate gene for mediating the development of both interspecific and intraspecific plumage variation and demonstrates the applicability of findings in domestic species to understanding avian plumage patterning more broadly.
Hansen, T. E.; Corpuz, R. L.; Simmonds, T.; Aldebron, C.; Mason, C.; Geib, S. M.; Sim, S. B.
Show abstract
The olive fruit fly, Bactrocera oleae (Rossi) (Diptera: Tephritidae), is a specialist of fruits of the genus Olea and is a major pest of commercial olives due to their adverse impacts to olive production. In support of genomic and physiological research of the olive fly, we sequenced, assembled, and annotated two independent genomes, one from a wild-collected male and one from a wild-collected female. The resulting genomes are highly contiguous, collinear, and complete, attesting to the accuracy and quality of both assemblies. In addition to the autosomes captured as single contigs, the X and Y chromosomes were also captured as evidenced by the X chromosome showing diploid coverage in the female assembly compared to haploid coverage in the male assembly and the Y chromosome being entirely absent from the female assembly. These assemblies represent the first full chromosome-level assembly for Olive fly. In addition, a complete genome assembly of a known obligate symbiont to the olive fly, Candidatus Erwinia dacicola, was fully captured. The Ca. E. dacicola we report here is the most contiguous to date, represented with a gapless chromosome and two separate gapless plasmids. These genome assemblies, along with bacterial symbiont assembly, provide foundational resources for future genetic and genomic research in support of its management as an agricultural pest.
Magallanes, M. E.; Lessa, E. P.
Show abstract
Abrothrix olivacea (Waterhouse, 1837), the olive grass mouse, is a widely distributed sigmodontine rodent that inhabits a broad range of environments, from the hyper arid deserts of southernmost Peru and northern Chile to the Patagonian steppe to the humid temperate rainforests of southern South America. Its extensive ecological breadth, coupled with physiological adaptations to water scarcity, makes it an ideal model for studying environmental responses and phenotypic plasticity. Here, we present the first de novo scaffold-level genome assembly of A. olivacea, generated from short-read DNA sequencing. The 2.25 Gb assembly achieved a scaffold N50 of 123 Mb and a BUSCO completeness score of 98.61%, indicating high sequence completeness. Genome annotation identified 21,476 protein-coding genes, providing a valuable resource for evolutionary, ecological, and functional genomics. As a case study, we used this reference genome to explore gene expression and genetic divergence in kidney tissue from individuals inhabiting contrasting environments: the southern Andean rainforest and the Patagonian steppe. By integrating single-cell transcriptomic data from Mus musculus, we performed cell type deconvolution, revealing environment-specific expression patterns linked to renal function. This new genomic resource opens avenues for investigating local adaptation, population structure, and conservation genetics in one of South Americas most ecologically versatile and widely distributed rodents.
Wright, J. J.; De Weerd, H.; Lees, A. C.; Shaw, K. J.; Griffiths, S. M.
Show abstract
The scaly-sided merganser, Mergus squamatus, is an Endangered piscivorous duck which has been declining since the late 1900s due to habitat loss, over-hunting, and climate change. Despite being a species of global conservation concern and subject to ex- and in-situ conservation efforts, genomic research has been limited, hindering our understanding of its population genetic status and evolutionary history. In this study, we present the first fully annotated, chromosome-level genome for the scaly-sided merganser, generated using Oxford Nanopore long reads, Illumina short reads, and Hi-C sequencing. The final assembly spans 1.1 Gb across 307 scaffolds, 64 of which are anchored into 35 chromosomes, covering 99.5% of the genome. The assembly shows high contiguity (N50 = 84.3Mb) and completeness, with a Benchmarking Universal Single-Copy Ortholog (BUSCO) score of 98%. Repeat sequences comprise 9.55% of the genome. Homology-based gene annotation identified [~]15,200 protein-coding genes. A complete 16,624 bp mitochondrial genome was also assembled and annotated. Synteny analysis revealed strong chromosomal conservation across the wider Anatidae family, with evidence of lineage-specific rearrangements. Pairwise Sequential Markovian Coalescence modelling indicates recent stability in the effective population size of the species, with past declines coinciding with Pleistocene glacial cycles. Our high-quality genome provides an essential resource for conservation genomic and evolutionary studies of the scaly-sided merganser, supporting ongoing efforts to manage and protect this threatened species.
Ward, C. M.; Onetto, C. A.; Borneman, A. R.
Show abstract
Fungal and bacterial symbiosis is an important adaptation that has occurred within many insect species, which usually results in the relaxation of selection across the symbiont genome. However, the evolutionary pressures and genomic consequences associated with this transition are not well understood. Pathogenic fungi of the genus Ophiocordyceps have undergone multiple, independent transitions from pathogen to associate, infecting soft-scale insects trans-generationally without killing them. To gain an understanding of the genomic adaptations underlying this transition, long-read sequencing was utilized to assemble the genomes of both Parthenolecanium corni and its Ophiocordyceps associate from a single insect. A highly contiguous haploid assembly was obtained for Part. corni, representing the first assembly from a single Coccoidea insect, in which 97% of its 227.8 Mb genome was contained within 24 contigs. Metagenomic-based binning produced a chromosome-level genome for Part. cornis Ophiocordyceps associate. The associate genome contained 524 gene loss events compared to free-living pathogenic Ophiocordyceps relatives, with predicted roles in hyphal growth, cell wall integrity, metabolism, gene regulation and toxin production. Contrasting patterns of selection were observed between the nuclear and mitochondrial genomes specific to the associate lineage. Intensified selection was most frequently observed across nuclear orthologs, while selection on mitochondrial genes was found to be relaxed. Furthermore, scans for diversifying selection identified associate specific selection within three adjacent enzymes catalyzing acetoacetates metabolism to acetyl-COA. This work provides insight into the adaptive landscape during the transition to an associate life history, along with a base for future research into the genomic mechanisms underpinning the evolution of Ophiocordyceps.